SEARCH RESULT

Year

Subject Area

Broadcast Area

Language

1 results listed

2025 Speech Emotion Recognition with Librosa using Support Vector Machines(SVM) and Convolutional Neural Networks (CNNs)

SER (Speech Emotion Recognition) is a notably advancing field. Its primary purpose is the identification of emotions delivered via speech; it is essential for applications involving computer and human interaction in fields such as healthcare, entertainment, and psychology. This research studies the use of SER in feature extraction implemented with the help of a library based on Python called Librosa, an application of CNN (Convolutional Neural Network) and SVM (Support Vector Machine) for segregating emotions. The issues are resolved by enhancing the reliability and precision of systems in the identification of emotions by analyzing the efficiency of CNN as well as SVM in segregating emotions such as Sadness, Surprise, Happiness, Neutral, and Anger. The outcomes are that CNN performs much more efficiently with 89.20% accuracy, whereas SVM only has about 80.50% accuracy. Consistency is maintained by CNN, which has higher F1 scores, recall, and precision in all categories. This proves its ability to deal with the complexities of segregating emotions delivered via speech. It is divulged through the confusion matrices that both models can give good performance while handling certain emotions, but CNN acquires a higher accuracy with fewer incorrect classifications, especially while figuring out emotions that have acoustic properties that are quite similar. The research concluded by stating that CNN is more suitable for tasks implementing SER because of its ability to capture detailed patterns of emotions in speech more efficiently than SVM. Future work may include architectures of deep learning, which are more advanced, like Transformer-based models or RNN, and having the dataset expanded so that there is an increase in generalization of different emotional backgrounds and expressions. Also, incorporating approaches to reduce noise and different audio environments would help enhance the model so that it adapts easily to applications used in the real world, offering applicable and robust SER systems.

International Conference on Advanced Technologies, Computer Engineering and Science
ICATCES

B. Pavin Dr.N.V. Chinnasamy

134 140
Subject Area: Computer Science Broadcast Area: International Type: Article Language: English